Papers with multi-modal encoder-decoder model

1 papers
Data-Efficient Playlist Captioning With Musical and Linguistic Knowledge (2022.emnlp-main)

Copied to clipboard

Challenge: Music streaming services feature billions of playlists created by users, professional editors or algorithms.
Approach: They propose a multi-modal encoder-decoder model for automatic playlist captioning that leverages linguistic and musical knowledge to generate correct and thematic captions.
Outcome: The proposed model yields 2x-3x higher BLEU@4 and CIDEr than state-of-the-art captioning algorithms on a new playlists dataset from two major streaming services.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations